Theoretical and Applied Genetics
○ Springer Science and Business Media LLC
Preprints posted in the last 90 days, ranked by how well they match Theoretical and Applied Genetics's content profile, based on 49 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit.
Sakurai, K.; Moreau, L.; Mary-Huard, T.; Charcosset, A.; Iwata, H.
Show abstract
In plant breeding, it is often necessary to improve a target trait while maintaining other essential traits within desirable ranges. When genetic relationships exist among these traits, improvements in the target trait may lead to undesirable changes in essential traits, complicating cross selections. In such cases, it is critical to select cross-pairs that are expected to produce progeny that satisfy the requirements for all traits. The progeny distribution of each crossing pair can be predicted using the estimated genotypic values and genetic (co)variances of the target and essential traits. By utilizing this distribution, the probability of generating progeny that satisfy predefined trait requirements can be evaluated, allowing a direct comparison of alternative crosses. In this study, we developed Cross Potential Selection for Multiple Traits (CPS-MT), a breeding strategy designed to improve a target trait while maintaining one or more essential traits within desirable ranges. CPS-MT extends the original Cross Potential Selection (CPS) framework to explicitly handle trade-offs between traits under genetic correlations. We evaluated the performance of CPS-MT through simulations involving four types of genetic relationships and two genetic causal factors between traits, resulting in seven scenarios. Across all scenarios, CPS-MT consistently improved the likelihood of obtaining desirable progeny, indicating that CPS-MT provides a practical and effective framework for cross selection under multi-trait constraints in breeding programs. Article SummaryThis study developed Cross Potential Selection for Multiple Traits (CPS-MT), a new breeding strategy designed to improve a target trait while maintaining one or more essential traits within desirable ranges. CPS-MT evaluates crossing pairs by predicting progeny distributions based on estimated genotypic values and genetic covariances, enabling direct comparison of alternative crosses under multi-trait constraints. Through simulations incorporating four types of genetic relationships and two causal factors (seven scenarios), CPS-MT consistently increased the likelihood of obtaining progeny that satisfied the predefined trait requirement. These results indicate that CPS-MT provides a practical, robust framework for target trait improvement under trait constraints.
Liu, D.; Zhang, X.; Snyman, L.; Garrard, T.; Wallwork, H.; Dadu, H.; Maclean, M.; Tong, J.; Chen, C.; Gamaralalage, D. J.; Periyannan, S.; Hickey, L.; Hayes, B. J.; Dinglasan, E.
Show abstract
Net blotch, caused by Pyrenophora teres, is a major constraint to barley production worldwide and occurs as two epidemiologically distinct forms: net form net blotch (NFNB) and spot form net blotch (SFNB). Although numerous resistance loci have been reported in recent years, their genetic relationship remains poorly understood, and the effective deployment of resistance is constrained by the complex genetic architecture of net blotch resistance. In this study, we used a haplotype-based mapping approach to dissect the genetic basis of resistance to NFNB and SFNB in a diverse panel of 950 barley accessions from the Australian Grains Genebank (AGG). Disease responses were evaluated across 13 experiments, and a total of 40 quantitative trait loci (QTL) were identified, including 26 associated with NFNB, 29 with SFNB, and 15 common for both diseases. Most loci co-localized with previously reported QTL, while six putative novel haploblocks highlighted untapped genetic diversity within the AGG collection. Correlation analyses across phenotypic, genetic and haploblock levels revealed a partial but incomplete overlap in resistance mechanisms between NFNB and SFNB. Among the 4,497 haploblocks, approximately 60% of them showed positive local genetic correlations between the two diseases, suggesting shared genomic contributions to resistance. Haplotype composition analysis further identified a resistant haplotype group, mainly comprising accessions of Asian origin, that exhibited high levels of resistance to both forms of net blotch. Through in-silico haplotype stacking simulations, we demonstrated the cumulative genetic potential achievable by combining favourable haplotypes. When the breeding objective was to improve resistance to both NFNB and SFNB, dual-disease stacking strategies outperformed single-disease approaches, highlighting the value of prioritising haplotypes with positive pleiotropic effects. Overall, this study provides a comprehensive haplotype-level framework for understanding net blotch resistance and delivers practical insights for breeding barley cultivars with durable and broad-spectrum resistance to both NFNB and SFNB.
Kadoumi, R.; Heslot, N.; Henriot, F.; Murigneux, A.; Berton, M.; Moreau, L.; Charcosset, A.
Show abstract
Modern hybrid maize (Zea mays L.) breeding programs are based on the management of distinct complementary heterotic groups to maximize heterosis in high-performing hybrids. This practice lowers shared genetic segments and increases divergence between groups to limit inbreeding in hybrids. However, most breeding programs have not always enforced strict separation between heterotic groups in the past. Competitor commercial hybrids were notably a common elite germplasm source for inbred development, which would diminish divergence between groups. This study proposes a new haplotype-based approach to assess hybrids residual inbreeding based on parental similarity. The new haplotype method has a stronger significant negative effect on hybrids grain yield than raw SNP data. Evaluation of modern experimental hybrids uncovered related inbreds contributing to superior rates of residual inbreeding. Analysis of these inbreds revealed haplotype transfers between heterotic groups, originating notably from the use of a Stiff Stalk-Iodent commercial hybrid as breeding starts material in both Stiff Stalk and Non-Stiff Stalk breeding populations. The introduction of this intergroup parent generated heterotic-group-specific haplotype migration between crossing pools. These fragments caused significant genome-wide residual inbreeding in experimental hybrids across selection cycles. This study highlights the necessity for accurate evaluation of external sources of diversity to minimize haplotype transfers and admixture between crossing pools. We demonstrate the consequences of using commercial hybrids in inbred development, particularly regarding residual inbreeding, and their effects on hybrid performance. Insights from these results can assist breeders in optimizing the choice of parents for introducing genetic diversity in a reciprocal recurrent selection scheme. KEY MESSAGEHaplotype-based hybrids parental similarity better predicts grain yield than marker-based identity-by-state. Utilization of commercial hybrids as breeding start material resulted in higher hybrid residual inbreeding even after several selection cycles
Hamazaki, K.; Tsuda, K.
Show abstract
Background: Germplasm collections contain wide genetic diversity that is valuable for plant breeding, but conducting phenotypic evaluation for all genotypes in field trials is rarely feasible. Bayesian optimization offers a way to decide, season by season, which genotypes to cultivate in order to identify superior genotypes with fewer evaluations. However, standard Bayesian optimization commonly starts from randomly selected genotypes and mainly relies on surrogate models built from marker genotype information, while the text-based passport information that accompanies germplasm is not fully used. We examined whether pre-trained large language models can provide prior knowledge that improves these decisions in germplasm evaluation. Results: We constructed a large-language-model-guided Bayesian optimization framework that introduces large language models into two parts of the Bayesian optimization workflow. In zero-shot warmstarting, a large language model proposes initial genotypes using passport information such as cultivar name, country of origin, and subpopulation, optionally together with principal component scores derived from genome-wide single-nucleotide-polymorphism markers. In addition, we evaluated a large-language-model-based surrogate model that predicts phenotypic values for untested genotypes using in-context learning from previously evaluated genotypes. Using a rice germplasm panel and two target traits (seed number per panicle for maximization and protein content for minimization), we compared strategies. For seed number per panicle, zero-shot warmstarting with a general-purpose instruction-following model reduced the number of evaluated genotypes needed to reach the best genotype, whereas improvements were small for protein content. When genomic information was available, Gaussian-process-based Bayesian optimization was the strongest overall approach, while the large-language-model-based surrogate model outperformed random baselines and was competitive in some settings. When genomic information was not available, predictions based on passport information improved efficiency compared with fully random strategies. Conclusions: Pre-trained large language models can inject useful agronomic knowledge into Bayesian optimization for germplasm evaluation, particularly by improving early-stage genotype selection, and can also support optimization when genomic information is unavailable. As models better handle long genomic sequences together with passport information, large-language-model-guided Bayesian optimization may become a practical and explainable decision-support approach for agricultural optimization.
de Freitas, G. M.; Certuche, D. S.; Jannink, J.-L.; de Oliveira, E. J.; Garcia, A. A. F.
Show abstract
Multi-trait genomic prediction offers a practical route to improve selection for costly, complex traits in clonally propagated crops such as cassava. In a Brazilian breeding panel of 1,078 cassava clones genotyped with 25,923 SNPs and phenotyped for six agronomic traits, we compared single-trait (ST) and multi-trait (MT) GBLUP models. Stage-wise mixed models produced BLUEs that fed into ST and MT-GBLUP. We tested five cross-validation schemes that mimic breeder realities: ST baseline (CV1); naive all-traits MT prediction for unphenotyped candidates (CV2); MT prediction using auxiliary trait phenotypes in the test set (CV3); and two sparse-phenotyping regimes with missingness by trait (CV4) or by clone (CV5) at 25%, 50%, and 75% levels. The main results were that, under the ST baseline (CV1), predictive ability ranged from 0.50 for DMC and 0.45 for FRY down to 0.13 for Le.Dis. A naive full MT model (CV2) performed approximately on par with ST-GBLUP. In contrast, MT designs (CV3) that included informative auxiliary traits, such as shoot yield and combinations with plant vigor and leaf disease severity, yielded small gains for DMC with predictive ability of approximately 0.51 (+2%), while FRY predictive ability increased to approximately 0.65 (+44%), accompanied by RMSE reductions for FRY up to approximately 13.5% (e.g. RMSE approximately 6.2). Sparse-phenotyping simulations (CV4/CV5) demonstrated that MT models sustain or even improve predictive ability under realistic missing-data regimes (PA {approx} 0.62 - 0.65). Selection concordance between MT and ST top-10% sets was generally high (>0.80), and MT configurations produced measurable improvements in expected selection response and genetic gain per cycle for several target traits. These results indicate that strategically implemented MT-GBLUP, using a small set of biologically and operationally informative auxiliary traits and optimized sparse phenotyping, can materially increase predictive accuracy and selection efciency for economically critical cassava traits while reducing phenotyping burden.
Sharma, S.; Gustin, J. L.; Frei, U. K.; Settles, A. M.; Lübberstedt, T.; Resende, M. F. R.; Hershberger, J.
Show abstract
Key messageA single-kernel near-infrared reflectance spectroscopy-based sorter can effectively identify haploid kernels for doubled haploid production in field and sweet corn backgrounds. Doubled haploid (DH) technology significantly shortens the breeding cycle for developing homozygous inbred lines in maize (Zea mays). Manual sorting of haploids from a larger bulk of hybrid kernels in an induction cross is a major bottleneck in DH development. Automated systems based on near-infrared (NIR) reflectance spectroscopy can be valuable tools for rapid haploid sorting, provided that sorting accuracy is sufficient for incorporation into the DH process. In this study, we evaluated the accuracy of a custom-built single-kernel NIR (skNIR) sorter for classifying haploid kernels from 12 high-oil haploid induction populations generated from two sweet corn and two field corn donors and four high-oil haploid inducers (HOHIs). We evaluated several general classification models that can be applied without population-specific recalibration or prior genotyping, including models that classified haploids based solely on predicted oil content, as well as multivariate methods that used all wavelengths of the NIR spectra. The highest classification accuracy was obtained using a general multivariate support vector machine (SVM) model. When combined with the two best-performing HOHIs, the general SVM model accurately sorted induction populations from two of the three donor backgrounds crossed with these inducers. Two oil-based methods showed less accurate classification than the multivariate SVM model, due to overlapping oil content distributions across the two kernel classes. Overall, this study demonstrates effective skNIR-based sorting of haploid kernels from diverse induction populations using a single general model. The practical deployment of this instrument in maize breeding programs is discussed.
Francisco, F. R.; de Oliveira, G. L.; Niederauer, G. F.; Fritsche-Neto, R.; Souza, A. P. d.; Furlan, M. F. M.
Show abstract
Although grapevine (Vitis spp.) is among the oldest and most economically significant fruit species globally, its genetic improvement faces major bottlenecks due to long juvenile periods and extended cycles for phenotypic evaluation. In this context, genomic selection (GS) has emerged as an effective alternative to traditional selection, offering a robust framework to optimize breeding programs by significantly reducing generation intervals while enhancing predictive accuracy (PA) in early generations and expected genetic gains (EGGs). Nevertheless, factors such as minor allele frequency (MAF) and population size can significantly affect predictive models, even to the point of making their use unfeasible in breeding programs. In this context, this study evaluated the effect of data dimensionality reduction on GS accuracy by selecting single-nucleotide polymorphisms (SNPs) based on MAF thresholds. The experimental design tested the predictive capacities of four machine learning (ML) algorithms (ElasticNet, K-Neighbors, Support Vector Machine Regression, and XGBoost) alongside the conventional Genomic Best Linear Unbiased Prediction (gBLUP) model. These were validated using three SNP datasets (11,115, 9,494, and 6,100 markers) filtered by MAF levels of 0.05, 0.1, and 0.2 across six genetic traits, and EGGs were compared between conventional breeding and GS via the breeders equation. The results revealed that the ML models exhibited remarkable stability, with no significant differences in PA across the different MAF-based SNP densities, except for berry length, which showed a substantial difference with XGBoost at an MAF of 0.2. Conversely, gBLUP demonstrated high sensitivity to dimensionality reduction, with its performance significantly impacted by MAF filtering across all the traits. These results suggest that compared with traditional GS models that rely on a genomic kinship matrix, ML-based approaches offer greater flexibility in feature reduction. Additionally, compared with chemical traits, morphological traits generally had greater predictive ability. Furthermore, every GS model provided estimated genetic gains superior to traditional breeding, with improvements ranging from an 8.90-fold increase in berry length to a 2.86-fold increase in total soluble solids, confirming that GS integration is promising for enhancing breeding efficiency in grapevines.
de Freitas, G. M.; Certuche, D. C. S.; Jannink, J.-L.; De Oliveira, E. J.; Garcia, A. A. F.
Show abstract
Genomic selection has become an important strategy in cassava breeding, enabling faster selection cycles and sustained genetic progress. Despite its widespread adoption, long-term evaluations integrating predictive performance, realized genetic gain, and genetic diversity remain scarce, particularly in clonally propagated crops. We present a comprehensive assessment of genomic selection outcomes in the Brazilian cassava breeding program across four recurrent selection cycles (C0 to C3) implemented between 2011 and 2024, using historical phenotypic and genomic data from 210 multi-environment trials. Predictive ability of genomic best linear unbiased prediction models ranged from low to moderate, depending on the traits genetic architecture and heritability. Prediction accuracies were highest in early cycles (C0 and C1) and showed modest declines in later cycles (C2 and C3). Root yield, shoot yield, plant height, starch content, and dry matter content exhibited stable predictive performance across cycles, with a gradual reduction in RMSE, indicating improved model calibration as training populations expanded. Regression analyses of genomic estimated breeding values revealed significant realized genetic gains for most yield-related traits. In contrast, dry matter content and starch content exhibited small, non-significant negative trends, consistent with known unfavorable genetic correlations with yield. Targeted reductions in plant architecture scores reflected deliberate selection for ideotypes suited to mechanized production systems. At the same time, analyses of genetic diversity revealed a slight decrease in observed heterozygosity, with higher values in the most advanced selection cycle. These results provide an integrated framework for monitoring predictive performance, realized genetic gain, and population genetic dynamics under long-term genomic selection. Collectively, they offer valuable insights into balancing short-term genetic improvement with long-term sustainability and support the development of strategies to optimize selection decisions, breeding planning, and population management in Brazilian cassava breeding programs.
Mas Gomez, J.; Rubio Angulo, M.; Duval, H.; Dicenta, F.; Martinez-Garcia, P. J.
Show abstract
In plant breeding and genetics, recent advances in high-throughput phenotyping are beginning to meet the growing demand for large-scale, high-quality phenotypic data that emerged after the development of next-generation sequencing technologies. Recent developments in phenomics have been incorporated into almond breeding programs, facilitating the large-scale acquisition of quantitative phenotypes and the dissection of the genetic architecture underlying morphological and quality-related traits. The implementation of a high-throughput phenotyping platform integrating RGB and hyperspectral imaging with genotyping using the 60K almond SNP array enabled the large-scale characterization of almond populations and the identification of 567 robust marker-trait associations across 66 traits. These analyses revealed two major genomic hotspots on chromosomes 2 and 5 associated with morphological and quality-related traits. These regions harbored biologically relevant candidate genes, including genes associated with OVATE family proteins, brassinosteroid signaling, protein ubiquitination, and acyl-CoA metabolism, as well as other regulators of organ growth, cell proliferation, hormone signaling, and seed development. Furthermore, a novel candidate gene encoding a COMT-like O-methyltransferase involved in lignin biosynthesis was identified and proposed to contribute to shell hardness, a major genetically controlled trait in almond. Together, these findings demonstrate the potential of integrating high-throughput phenomics and genomics to dissect complex traits, identify candidate genes, and accelerate genomics-informed breeding in almond.
Tajima, A. M.; Matthews, W. C.; Duong, T.; Khanh, T. D.; Baniya, A.; Penmetsa, R. V.; Parker, T.; Farmer, A.; English, S.; Diepenbrock, C.; Gepts, P.; Roberts, P. A.; Huynh, B.-L.
Show abstract
Lima bean (Phaseolus lunatus) is a broadly adapted, economically important leguminous crop and a susceptible host of root-knot nematodes (Meloidogyne spp.; RKN), which are a devastating plant pathogen in agricultural systems worldwide. To date, there have been few studies to elucidate the genetic determinants of RKN resistance in lima beans. Understanding the genetic mechanisms underlying resistance is essential for improving resistance traits and incorporating them into lima bean breeding programs. To assist in marker-assisted selection, we aimed to identify and map quantitative trait loci (QTLs) conferring RKN resistance-related traits. Three recombinant inbred line (RIL) populations were used in this study. Three populations were derived by crossing two RKN-resistant parents with the same RKN-susceptible parent and with each other. All populations were genotyped using genome-wide single-nucleotide polymorphism (SNP) markers. Each population was screened for root galling (RG) and RKN egg reproduction (ER) in response to M. incognita and M. javanica in greenhouse experiments. Three major QTLs were detected and mapped on chromosome Pl04 (QRk-pl04.1), Pl05 (QRk-pl05.1) and Pl10 (QRk-pl10.1) across populations. Among them, QRk-pl05.1 and QRk-pl10.1 affected levels of RG and ER of both RKN species, while QRk-pl04.1 suppressed root galling and reproduction responses of M. incognita but not of M. javanica. These chromosomal regions defined by flanking markers will help guide marker-assisted breeding and gene discovery for broad-based RKN resistance in lima beans.
Guffanti, F.; Nagel, K. A.; Galinski, A.; Mueller, C.; Pariyar, S. R.; Scheuermann, D.; Urbany, C.; Presterl, T.; Ouzunova, M.; Schoen, C.-C.
Show abstract
Characterizing the genetic basis of root system architecture and its role in early plant development is essential for developing maize varieties with improved nutrient uptake, enhanced early vigour, and higher yield potential in temperate regions. Landraces represent an invaluable source of allelic diversity that can be leveraged to enrich the genetic basis of modern breeding material. In this study, we used a high throughput phenotyping platform to characterize genetic variation for seedling root traits under chilling conditions relevant for early plant establishment in a large doubled haploid (DH) library derived from two European maize landraces. We dissected the quantitative genetic architecture of twelve seedling root traits using a haplotype-based genome-wide association study, identifying large-effect haplotypes specific to the individual landraces as well as numerous small-effect haplotypes present in both landraces. We validated the effects of four QTL in a biparental population, demonstrating their stability across genetic backgrounds. We found highly significant correlations between haplotype effects on seedling root traits evaluated in the phenotyping platform and early plant height evaluated in multi environment field trials, demonstrating the relevance of seedling root architecture for early plant establishment. In particular, haplotypes associated with seminal and lateral root length were the major determinants of early plant height under field conditions. Several of the haplotypes increasing seedling root length were absent from a broad panel of flint breeding lines, highlighting their potential as targets for introgression to improve early plant establishment under temperate growing conditions. Key messageSeedling root QTL discovered in a high throughput phenotyping platform under chilling conditions influence early plant development in the field.
Daware, A. v.; Hacke, C.; Remay, A.; Starnberger, P.; Schraml, C.; Collonnier, C.; Laurens, F.; Schmid, K. J.
Show abstract
Testing for distinctness, uniformity, and stability (DUS) is a requirement for plant variety registration and based on phenotypic traits, which is time-consuming and sensitive to environmental variation. Advances in genomics allow to complement DUS testing with molecular markers, for which two models in DUS testing were proposed by the Union for the Protection of New Varieties of Plants (UPOV). A use cases was described for maize, but an implementation has been hindered by a lack of suitable markers and validated analytical frameworks. We address these challenges by integrating historical DUS characteristics scores from 352 European hybrid maize varieties with high-density genome-wide single nucleotide polymorphism (SNP) data. Using genome-wide association studies (GWAS), we identified 18 genomic regions and candidate genes associated with 12 DUS characteristics, enabling the development of diagnostic markers consistent with the UPOV model "Characteristic-Specific Molecular Markers". Since most DUS traits are polygenic, we combined GWAS-informed marker selection with XG-Boost-based machine learning to predict notes of DUS characteristics. This approach achieved strong predictive performance across multiple traits (mean accuracy 0.67), demonstrating its potential for managing reference collections under UPOV model "Combining phenotypic and molecular distances in the management of variety collections". Both approaches were validated for two characteristics using independent public USDA-NPGS maize datasets (>1,700 accessions) highlighting the value of public data for method validation. We also identify key limitations of historical DUS data, including imbalanced and sparse trait representation, and discuss mitigation strategies. Despite these constraints, our results demonstrate that molecular markers may improve maize DUS testing, enabling faster, more accurate variety registration and supporting accelerated crop improvement. Key messageHistorical DUS datasets can be used to identify marker-trait associations of DUS characteristics using genome-wide association study (GWAS) and to develop a genomic prediction framework for an accurate prediction of DUS character notes from marker data.
Baraja-Fonseca, V.; Gil-Villar, D.; Bancic, J.; Renau-Morata, B.; Salud Justamante, M.; Plazas, M.; Gramazio, P.; Vilanova, S.; Perez-Perez, J. M.; Granell, A.; Molina, R. V.; Nebauer, S. G.; Prohens, J.; Arrones, A.
Show abstract
Nitrogen-use efficiency (NUE) is a pivotal breeding target in tomato (Solanum lycopersicum L.) to sustain production under reduced N inputs. Here, we leveraged a recently developed tomato multi-parent advanced generation inter-cross (ToMAGIC) population to identify lines with superior performance under reduced N availability. The eight founders and a core subset of 118 ToMAGIC lines were characterized with 10,684 SNP markers and evaluated under optimal (opN, 15 mM) and suboptimal (subN, 8 mM) N supply in an experiment totalling 1,576 plants, generating 48,068 data points across 61 phenotypic variables. Under both N treatments, ToMAGIC lines exhibited transgressive segregation for most traits, confirming the value of this population as a reservoir of untapped variation. Notably, under subN conditions, harvest index (Hi) increased by 29-44%, suggesting adaptive resource redistribution toward reproductive sinks. Variance partitioning revealed that agronomic and NUE-related traits were largely under genetic control, with heritability estimates frequently above 0.80 and broadly conserved across N treatments. Multivariate trait analysis identified fruit yield N concentration (NUE component, CN,y), shoot biomass N content (NAb), and shoot growth-related traits as the main drivers of treatment differentiation. Finally, proxy traits were prioritized by integrating response magnitude, heritability, trait correlations, and treatment-discriminatory power into multi-trait selection indices. This strategy generated favorable predicted genetic gains, reaching 158% for high-performance lines and 170% for subN-adapted lines, and consistently identified lines 402, 428, 518, 800, and 816 as promising pre-breeding materials. Overall, this study supports ToMAGIC as a powerful resource for developing N-efficient cultivars suited for sustainable agriculture.
Solarte Certuche, D. C.; Mamedio de Freitas, G.; Jannink, J.-L.; Garcia Morales, C. F.; Sousa Cerqueira, T.; Santos de Santana, B.; Jorge de Oliveira, E.; Garcia, A. A. F.
Show abstract
Post-harvest physiological deterioration (PPD) represents a significant challenge of cassava production and commercialization. This multifaceted biological process involves a series of mechanisms, including enzymatic stress responses, alterations in gene expression, protein synthesis, accumulation of secondary metabolites, and ultimately, programmed cell death. These changes render the storage roots unpalatable and unmarketable. Therefore, unraveling the genetic architecture of PPD and exploring the interactions of associated genes during its early and late stages is essential for the crop production. We used modern genetic resources to unravel the genetic basis of PPD, based on a genome-wide association study (GWAS), utilizing a combination of different models, including BLINK (Bayesian-information and Linkage-disequilibrium Iteratively Nested Keyway), SUPER (Settlement of MLM Under Progressively Exclusive Relationship), and MLMM (Multi-locus mixed models). The phenotyping dataset spanned five years and included evaluations from 42 different trials, evaluating the Embrapa (Brazilian Agricultural Research Corporation) germplasm along with a population derivative from a genomic selection cycle. We utilized a genotype dataset comprising 26,000 high-quality SNPs (single nucleotide polymorphisms). Our findings indicated four significant genetic variants located on chromosomes 2, 5, and 13, which together explain 35.83 % of the phenotypic variation. These variants are associated with genes that are closely linked to the pathways activated during the early and late symptoms of PPD. The identification of these three key genes provides valuable insights into the genetic architecture of PPD and lays a strong foundation for molecular breeding, supporting the efforts to identify cassava genotypes with enhanced PPD tolerance, the identified genomic regions may be incorporated into genomic selection models, thereby enhancing marker-assisted selection (MAS) and improving breeding strategies for long shelf life and high-quality agronomic cassava cultivars for the cassava community.
Shaffer, W.; Papin, V.; Carter, Z.; Brunner, S. M.; Tong, J.; Villiers, K.; Robinson, H.; Voss-Fels, K.; Hayes, B. J.; Hickey, L.; Dinglasan, E.
Show abstract
Haplotype-based breeding strategies have emerged as promising approaches to maximize long-term genetic gain by identifying complementary parental combinations while maintaining genetic diversity. However, these methods typically require phased genotypes and more intensive workflow pipelines and skillsets. We developed a novel local genomic estimated breeding value (localGEBV) fitness function with similar intent to the optimal haplotype stacking (OHS) framework fitness function and implemented both in the novel R package, HapSelect. Our aim was to evaluate whether phased haplotypes provide additional benefit over the more easily available dosage-based unphased genotypes in highly inbred crops. A subset of bread wheat nested association mapping (NAM) population comprising 444 lines genotyped with 6,054 DArT-Seq markers was analysed. Marker effects were estimated using rrBLUP, localGEBV and haplotype effects were calculated across linkage disequilibrium-defined haploblocks, and genetic algorithms (GA) were used to identify optimal sets of 30 founders using either a localGEBV derived fitness function with unphased, dosage inputs or the OHS fitness function with phased inputs. Selected parental sets were compared with conventional truncation selection (TS) through 150 generations of forward simulation. The OHS fitness function achieved a marginally greater optimized ultimate GEBV than the localGEBV fitness function during GA optimization, with only 18 of the 30 selected founders overlapped between the two methods. Despite these differences, forward simulations demonstrated nearly identical long-term genetic gain for localGEBV and OHS-selected founders, with both approaches outperforming conventional truncation selection by maintaining greater genetic diversity and delaying the genetic plateau. The minimal difference between localGEBV and OHS is likely attributable to the high homozygosity of the population, where localGEBV and haplotype effects are nearly confounded. These results demonstrate that dosage-based localGEBV provides a practical alternative to phased haplotype approaches for parent selection in inbred crops, substantially simplifying genomic workflows while maintaining long-term breeding performance. Future work should evaluate these methods in more diverse inbred populations and outbred species, where great haplotypic diversity may increase the advantage of true haplotype-based optimizations.
Solarte Certuche, D. C.; Mamedio de Freitas, G.; Jannink, J.-L.; Garcia Morales, C. F.; Sousa Cerqueira, T.; Santos de Santana, B.; Jorge de Oliveira, E.; Garcia, A. A. F.
Show abstract
Cassava is a major staple crop in tropical regions, and improving its root nutritional quality, particularly carotenoid and dry matter content (DMC), remains a central breeding goal. To elucidate the genetic basis of these traits by locating genomic regions associated with them, we analyzed 3,043 cassava clones from the Brazilian Agricultural Research Corporation (Embrapa) breeding program, phenotyped across 188 multi-environment trials conducted from 2011 to 2022 in Brazil. All clones were genotyped using Genotyping-by-Sequencing (27,045 Single Nucleotide Polymorphism - SNPs) and Diversity Arrays Technology - DArTseq (25,923 SNPs). Trait values were estimated using a two-stage mixed model to obtain deregressed BLUPs (Best Linear Unbiased Predictions), and genome-wide association analyses were performed using both the Mixed Linear Model (MLM) and Multi-Locus Mixed Model (MLMM). We detected six significant SNPs consistently associated with carotenoid content and DMC after Bonferroni correction. These SNPs mapped to six candidate genes involved in pathways relevant to root physiology, including Abscisic Acid ABA-related signaling, hydrolase activity affecting carotenoid conversion, fatty-acid biosynthesis within plastids, cell-wall remodeling, and glycolytic energy metabolism. The loci jointly explained 75.56 % of the phenotypic variance for carotenoids and 76.23 % for DMC, with individual SNP effects ranging from [~]17 % to [~]42 % PVE (Proportion of Variance Explained). Broad-sense heritability was H2 = 0.78 for carotenoids and H{superscript 2} = 0.34 for DMC, confirming substantial genetic control and suitability for molecular breeding. Haplotype analyses revealed four superior haplotypes for carotenoids and one key haplotype for DMC, each showing significantly higher trait values compared with other allelic combinations. These haplotypes represent promising targets for marker-assisted selection and genomic selection, with direct applicability for accelerating genetic gain in elite breeding populations. The results provide actionable genomic resources for breeding programs aiming to develop biofortified and high-root quality cultivars and establish a foundation for future multi-omics and functional validation studies.
Rosati, C.; Tirado, F.; Aprea, G.; Bitton, F.; Brault, M.; duboscq, R.; Ferrante, P.; Pellegrino, K.; Stamigna, C.; Giuliano, G.; Causse, M.
Show abstract
Reducing postharvest fruit loss without compromising fruit quality is a major goal in tomato breeding. Fruit shelf life is a complex trait, influenced by postharvest changes in fruit firmness and weight, as well as by both genetic and environmental factors. A QTL mapping experiment was conducted to identify loci associated with fruit shelf life-related traits using two distinct F2 tomato populations. Fruit weight, firmness, shelf life (evaluated as both loss of fruit weight and loss of firmness over time), and colour space components were measured, and QTLs were mapped using a commercial low-density SNP genotyping panel and bulk segregant sequencing analysis experiments. We show that fruit firmness at harvest is only weakly predictive of postharvest firmness loss, indicating that shelf life should be treated as a dynamic trait rather than a static firmness phenotype. Across the two populations, 60 QTLs defining 26 genomic regions were identified, including both population-specific loci and shared regions on chromosomes 9 and 12. Several narrow intervals contained candidate genes related to ethylene signaling, cell-wall remodeling, calcium transport, aquaporin-mediated water balance and stress responses. The identified QTLs and candidate genes are directly relevant to breeding programmes seeking to improve postharvest performance without using major ripening mutants that compromise fruit quality. Key messageComparative QTL mapping and BSA-Seq in two F2 tomato populations show that postharvest shelf life is only partly explained by fruit firmness at harvest and identify candidate loci for firmness loss and weight loss during storage.
Hatta, T.; Hamazaki, K.; Fuji, Y.; Toda, Y.; Ichihashi, Y.; Ohmori, Y.; Yamasaki, Y.; Takahashi, H.; Takanashi, H.; Tsuda, M.; Tsujimoto, H.; Kaga, A.; Nakazono, M.; Fujiwara, T.; Hirai, M. Y.; Iwata, H.
Show abstract
Metabolic phenotypes are often governed by complex genetic architectures involving both additive and non-additive effects. However, the extent to which epistatic interactions contribute to the pathway-level regulation of plant metabolism remains unclear. In this study, we investigated the genetic architecture of flavonoid-related metabolites using metabolomic and genomic data from 200 soybean accessions cultivated under multiple environmental conditions. Broad-sense heritability estimates revealed that many metabolites were under strong genetic control, particularly flavonoid-related metabolites. Principal component analysis-based metabolome-wide genome-wide association studies identified four major loci associated with flavonoid metabolic variation, including a locus corresponding to flavonoid 3'-hydroxylase. Conditional analyses based on multilocus genetic backgrounds demonstrated that the effects of downstream loci were highly dependent on upstream genotypes. In particular, single-nucleotide polymorphism effects were frequently detectable only in specific allelic backgrounds defined by the major flavonoid 3'-hydroxylase locus, consistent with strong epistatic interactions among loci. Bayesian network analyses further supported a hierarchical genetic structure consistent with upstream regulation of downstream loci across the flavonoid biosynthetic pathway. These results demonstrate that highly heritable metabolic phenotypes can be controlled by a few loci exhibiting both additive and context-dependent non-additive effects. Our findings provide evidence that pathway-level metabolic diversity in soybean is generated through hierarchical and epistatic genetic control involving a limited set of key loci.
Harris, Z. N.; Braley, J.; Cassetta, E.; Crain, J.; DeHaan, L.; Van Tassel, D.; Miller, A.; Rubin, M. J.
Show abstract
Perennial grains represent a promising frontier for sustainable agriculture, but breeding progress is constrained by the accessibility of genotyping and the difficulty of evaluating complex traits expressed for multiple years after establishment across heterogeneous environments. Phenomic selection may help address these challenges by using inexpensive, scalable, high-dimensional phenotypes collected early in development, although the robustness of such predictions across breeding cycles remains uncertain. Here, we compared genomic selection and phenomic selection across two breeding cycles of Thinopyrum intermedium (intermediate wheatgrass; IWG; Kernza(R)), comprising approximately 2,280 individuals from maternal half-sib families evaluated across multiple field sites and years. We constructed relationship matrices from genomic markers and early-life stage phenomic data, including seed and leaf color (HSV), CropReporter multispectral reflectance and indices, and cycle-specific hyperspectral reflectance sensors. Genomic models provided the strongest predictions on average across all field traits in both cycles. Among phenomic predictors, leaf HSV was consistently the most informative, whereas CropReporter and hyperspectral data showed lower and more trait-dependent performance and seed HSV provided little predictive value. Genomic, leaf HSV, and CropReporter models transferred across breeding cycles with little apparent loss of predictive ability relative to within-cycle validation, demonstrating that their predictive signals were not restricted to a single breeding cycle. Early-life stage leaf HSV emerged as a practical, accessible tool for germplasm thinning and early-stage prioritization in perennial breeding programs. Despite limited similarity among relationship matrices, multi-relationship-matrix models rarely improved prediction beyond the stronger constituent single-relationship-matrix model. Together, these results show that early-life stage phenomic data provide reproducible information about agronomic performance expressed years later, but that predictor complexity and data integration do not guarantee improved prediction.
Degen, B.; Leite Montalvao, A. P.; Schneck, V.
Show abstract
Long breeding cycles constrain genetic gain in Scots pine, while mature progeny trials can provide reference populations for genomic selection. We combined phenotypic records from two German trials established in 1990 with dense SNP data. We complemented these data with an offspring-level parentage audit, duplicate filtering, and simulations of breeding strategies. Of 1,986 phenotype-matched genotyped trees, 1,668 formed the parentage-green set and 1,651 remained after exclusion of 17 near-duplicate samples. The final population comprised 1,521 supported controlled-cross offspring and 130 mother-known, father-unknown offspring from 87 progeny labels. Across 14 trait-by-age measurements, GBLUP gave the clearest improvements for diameter and volume from age 20 onwards, whereas height was more mixed. At age 35, ABLUP versus GBLUP heritability was 0.155 versus 0.235 for diameter, 0.265 versus 0.263 for height, and 0.172 versus 0.236 for volume. Leave-progeny-out predictive ability increased from 0.205 to 0.238, 0.249 to 0.265, and 0.203 to 0.239, respectively. SNPscan_breeder simulations compared phenotypic selection, progeny testing, cross-generation genomic selection, and genomic selection with phenotypic thinning under five diversity variants. Progeny testing produced the greatest cumulative gain, but its 45-year cycle reduced annual response. Under the base assumptions, genomic selection with phenotypic thinning gave the highest annual gains for height and fungal resistance, whereas pure genomic selection gave the highest annual diameter gain. Genomic strategies accumulated more kinship than conventional strategies, although a {lambda} = 0.10 kinship penalty improved founder retention with little loss of gain; no mitigation option simultaneously maximised gain, prediction accuracy, and diversity. These results support genomic shortlisting within a field-tested programme with explicit reference-population updating and diversity management.